Microbial Genomics
● Microbiology Society
Preprints posted in the last 90 days, ranked by how well they match Microbial Genomics's content profile, based on 225 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit.
Ashcroft, M.; McGarry, N.; Stanton, T. D.; Hoyles, L.; Holt, K. E.; Wyres, K. L.
Show abstract
The Klebsiella oxytoca Species Complex (SC) represents an emerging healthcare-associated group of opportunistic pathogens. The capsular polysaccharide is a virulence determinant and target for novel vaccines, monoclonal antibodies and phage therapy. In the absence of broadly accessible phenotyping techniques, prediction of capsule types from whole-genome sequence data is critical for understanding capsule diversity and epidemiology, and to prioritise capsule types as targets for novel anti-K. oxytoca SC interventions. Here we present the first comprehensive capsule synthesis locus (K locus) database targeted for the K. oxytoca SC, comprising 88 distinct loci defined by gene content and which is compatible with the rapid genome typing tool Kaptive. The database provides high coverage of publicly available K. oxytoca SC genomes (97.6% of 2,244 genomes, dereplicated from a total of 4,055), and the typing rate is significantly higher than that achieved with the pre-existing Klebsiella K locus database (97.6% vs 50.3%, p <0.0001), which primarily targets the Klebsiella pneumoniae SC. We demonstrate the utility of the novel K. oxytoca SC database by application to three diverse clinical K. oxytoca SC isolate collections (n=61 to 102 genomes each), suggesting a high diversity of K types. The novel K. oxytoca SC K locus database (github.com/klebgenomics/KoSC-surface-antigen-loci) will provide a key resource to support larger systematic studies and ongoing genomics surveillance efforts for the K. oxytoca SC. IMPACT STATEMENTMembers of the Klebsiella oxytoca Species Complex (SC) are an emerging cause of infections in humans and are frequently associated with antimicrobial resistance. Klebsiella species produce two key surface antigen sugars (capsular polysaccharide and lipopolysaccharide) that are immunogenic and are targets for novel control strategies such as vaccines and phage therapy. Phenotypic typing of these surface antigen sugars (serotyping) is costly and laborious, with genotyping (predicting the serotype from whole-genome sequence data) a useful alternative. Here, we present a curated capsule (K) locus reference database for the K. oxytoca SC, which represents a useful tool to assist in the epidemiological surveillance of this emerging pathogen. DATA SUMMARYAll Klebsiella oxytoca Species Complex genomes used in this work were publicly available, with accession details listed in Supplementary Tables 1, 2 and 3. The K. oxytoca Species Complex K locus reference database is available under GNU public license at github.com/klebgenomics/KoSC-surface-antigen-loci.
Parker, M. J.; Hopkins, K. M. V.; Chau, K. K.; Cregan, J.; Oakley, S.; Barrett, L.; Jeffery, K.; Butcher, L.; Paulus, S.; Young, B. C.; Eyre, D. W.; Fowler, P. W.; Stoesser, N.; Sanderson, N. D.; Bejon, P.
Show abstract
Antimicrobial resistance genes (ARGs) can spread via horizontal transfer or clonal expansion. We investigated the genomic epidemiology of extended-spectrum beta-lactamase (ESBL)--producing Klebsiella pneumoniae and Escherichia coli in a neonatal unit. Between January and November 2023, 53 ESBL isolates were obtained from 23 neonates via routine screening and clinical sampling. Long-read nanopore sequencing identified blaCTX-M-15 as the dominant ESBL gene, alongside blaCTX-M-65 and blaCTX-M-27. Among 49 blaCTX-M-15 isolates, 33 carried the gene on plasmids and in the remainder it was located on the chromosome. ESBL isolates belonged to one of seven MLST sequence types; E. coli isolates were dominated by ST131/ST131-like/ST5640 lineages, while K. pneumoniae were predominantly ST13. Plasmids (ESBL and non-ESBL-associated) clustered into 18 communities, five of which contained plasmid-bearing blaCTX-M genes. The largest cluster comprised IncFIB(K) plasmids from K. pneumoniae ST13, although these were predicted as "non-mobilizable" by MOB-suite and belonged to isolates from a clonally disseminated strain. No evidence of blaCTX-M dissemination via shared plasmids was identified. Meanwhile, chromosomal phylogenetic analysis identified four distinct clonal clusters with [≤]7 SNP differences (n = 10, 5, 5, and 2 patients). In this setting, genomic analysis supported clonal dissemination of several blaCTX-M-associated strains as the main outbreak mechanism, affecting 20/23 neonates, rather than plasmid-mediated transmission.
Makaranga, A.; Bahati, S. Y.; Mwakalapa, E. B.; Mung'ong'o, H.; Kalolo, A.; Thomas, C.; Kassam, D.; Chisembe, P.; Shibayama, K.; Maghembe, R. S.
Show abstract
Campylobacter jejuni is a leading foodborne cause of gastroenteritis, but genomic surveillance remains uneven across Africa. A recent East Africa study combined whole-genome sequencing and antimicrobial susceptibility testing for Campylobacter isolates from humans with diarrhea in Kenya and poultry in Tanzania, showing high sequence-type diversity and substantially higher multidrug resistance in poultry. We extended this regional evidence by analyzing 1,013 publicly available C. jejuni genomes, including 718 African and 295 non-African comparator genomes, with standardized assembly, genotyping, phylogenomics, pangenome reconstruction, antimicrobial resistance, virulence, and mobile-element profiling. African genomes were geographically concentrated but genetically diverse, included globally distributed and regionally enriched lineages, and showed an open pangenome dominated by low-frequency gene families. Resistance and virulence determinants were unevenly distributed by region and lineage. These findings place African C. jejuni diversity within a global evolutionary framework and support expanded, integrated One Health genomic surveillance. Data summaryAll genome sequence data analysed in this study were retrieved from publicly accessible repositories, including the National Center for Biotechnology Information Sequence Read Archive and Assembly resources and corresponding records available through the International Nucleotide Sequence Database Collaboration where applicable. Accession identifiers, BioSample records, run accessions, country metadata and host/source information for all analysed genomes are provided in the combined Supplementary Data workbook. No new sequence data were generated. The analysis used publicly available data generated by other investigators, and the original data-generating studies are cited where appropriate. Derived analytical outputs supporting the findings are included in the manuscript and Supplementary Information. Impact statementGenomic surveillance of Campylobacter jejuni remains uneven globally, and African data are still underrepresented in many comparative analyses. This study brings together publicly available African and non-African C. jejuni genomes in a single standardized comparative framework, linking population structure, pangenome composition, antimicrobial-resistance determinants, virulence-associated loci and mobile-element profiles. The work shows that African C. jejuni diversity is not peripheral to the global population: African genomes include globally distributed sequence types, regionally enriched lineages and a large accessory-gene repertoire. By separating genomic surveillance signals from population-representative prevalence claims, the study provides a cautious framework for interpreting public genome collections from settings with unequal sampling. The findings support broader One Health genomic surveillance, improved metadata completeness and geographically balanced sequencing to better understand foodborne transmission, resistance evolution and lineage diversification in this important zoonotic pathogen.
Garcia Gonzalez, N.; Ferragud, R.; Blane, B.; Kim, J. I.; Torok, M. E.; Harrison, E. M.; Gouliouris, T.; Coll, F.
Show abstract
BackgroundGenomic prediction of antimicrobial resistance (AMR) relies on the accurate detection of resistance genes or allelic variants of core genes from raw or assembled genomes sequences. For several bacterial species and antibiotics, AMR genotype-phenotype discrepancies are common, indicating that important sources of error remain unresolved. For Enterococcus faecium, we focused on identifying the sources of discrepancies for tetracycline resistance, for which genotypic detection had shown particularly low accuracy. We investigated the effect of structural variation in antibiotic resistance genes (ARGs)--including gene duplications, truncations, interruptions, and mixed configurations of complete and partial gene copies-- as a source of genotype-phenotype discrepancies from short{square}read data. We conduct further extended investigations to other antibiotic families and into another bacterial species: Escherichia coli. MethodsWe analyzed collections of E. faecium and E. coli genomes, integrating high{square}quality complete assemblies, simulated Illumina short reads, and matched AMR phenotypic data. The integrity, copy number, and allelic diversity of ARGs were examined for multiple antibiotic classes, and their impact on ARG detection and accuracy of AMR determination was assessed using several commonly used bioinformatic tools (SRST2, ARIBA and AMRFinderPlus). ResultsFor E. faecium, after ruling out the effect of specific tet allelic variants on tetracycline susceptibility, we found that the integrity and copy number of tet(M) had a major effect on detection accuracy. Duplicated and incomplete ARGs are also common in E. faecium genomes, particularly for macrolides (erm(B)) and aminoglycosides (ant(6)-Ia and aph(3)-IIIa). In E. coli, similar patterns were observed for tet(A), erm(B) and aminoglycoside{square}associated genes (aph(3{square})-IIIa and ant(6)-Ia). Across ARGs in both species, short-read mapping methods wrongly reported interrupted genes as complete in some instances, while assembly{square}based methods often failed to resolve complete copies of duplicated genes. Detection accuracy improved when tools were adapted to account for gene integrity and when extended AMR databases incorporating species{square}specific alleles were included. ConclusionsOur findings reveal that bioinformatic limitations in dealing with ARG copy number and completeness, and in accounting for allelic variation, underly a substantial source of genotype-phenotype errors, highlighting the need for improved AMR databases and bioinformatic tools that consider these factors to achieve reliable genomic prediction of AMR.
Ryan, Y.; Jolley, K. A.; Hearn, H.; Parfitt, K. M.; Platt, S.; Lamagni, T.; Moganeradj, K.
Show abstract
Streptococcus pyogenes is a globally important pathogen responsible for at least 500,000 deaths a year, causing significant burden on healthcare systems. It is the causative agent for ailments such as impetigo and strep throat to septicaemia and necrotizing fasciitis. Assessment of genetic relatedness for the detection of outbreaks within communities or healthcare facilities is vital in decreasing the propagation of S. pyogenes within these settings, alongside epidemiological data. As the volume of isolates being sequenced increases year on year, more scalable and sharable methodologies of assessing genetic relatedness are required by reference laboratories and for international collaboration. LIN codes, applied to core genome MLST (cgMLST) represent a method which is extensible to large scale whole genome sequencing (WGS) while still being sufficiently sensitive to detect outbreak clusters. Here we present a novel cgMLST and LIN code scheme, hosted by PubMLST, enabling international collaboration and global tracking of variants, that is highly scalable and usable for all. The schemes are available at https://pubmlst.org/organisms/streptococcus-pyogenes. Data SummaryGenome sequences and metadata are available at https://pubmlst.org/organisms/streptococcus-pyogenes. PubMLST and ENA accessions and metadata can additionally be found in the supplementary data. Raw reads for UKHSA sequences are available in ENA study PRJEB115996. Impact StatementStreptococcus pyogenes is a globally relevant pathogen capable of causing invasive and non-invasive disease across a multitude of settings. Assessment of genetic relatedness is an increasingly important aspect of managing outbreaks, requiring solutions that are scalable, high resolution and comparable across laboratories. Here we present a high resolution core genome multi locus sequence typing (MLST) and associated life identification number (LIN) code scheme, The schemes were developed using a combination of 4,916 UKHSA and 2,391 publicly available S. pyogenes isolates in order to cover a wide range of EMM types both within the UK and globally. These new schemes enable high resolution typing of S. pyogenes isolates, suitable for analysis of lineages to genomic epidemiology in outbreak detection and management. Both cgMLST and LIN code schemes are available on PubMLST as an open access resource for the public health and academic communities and can enable both intra laboratory and global coordination.
Lubwama, M.; Hoyles, L.; McCartney, A. L.; Kateete, D. P.; Bwanga, F.; Kigozi, E.; Kalema, L.; Asiimwe, B.; Katende, G.; Lwigale, F.; Sekyanzi, S.; Niyonzima, N.; Orem, J.; Ddungu, H.; Kambugu, J.; Phipps, W.; Winter, J.
Show abstract
Antimicrobial resistance (AMR) exacerbates bacteraemia in cancer patients, particularly in low-resource settings. At the Uganda Cancer Institute, high rates of Enterobacterales producing extended-spectrum {beta}-lactamases (ESBLs) have been reported, with DNA-based detection of bla genes limited to PCR. This study aimed to determine whether bacterial genomic DNA shipped at ambient temperature from Uganda to the UK retained sufficient quality for whole-genome sequencing (WGS), to allow in-depth genomic analyses of isolates. Genomic DNA was extracted from Gram-negative bloodstream isolates (n=77) in Uganda and shipped to the UK at ambient temperature. rpoB gene (77/77, 100%) and WGS data (72/77, 93.5%) were generated for isolates, with 66/72 (91.7%) genomes of high-quality (Escherichia coli n=34; Klebsiella spp. n=32). Bioinformatic analyses included species identification, sequence typing, SNP analysis, AMR and virulence gene profiling, and comparison with publicly available genomes of Ugandan isolates. Phenotypic-genotypic concordance was generally high: 7/77 (9.1%) isolates were misidentified by phenotypic testing, and two showed unexplained carbapenem resistance. E. coli isolates showed diverse sequence types, with high prevalence of blaCTX-M (91.2%) and blaOXA-1 (47.1%); carbapenemase genes were rare. Klebsiella isolates lacked hypermucoidy loci and displayed diverse capsule types, with a high prevalence of ESBLs. Genomic clustering suggested limited within-hospital transmission of strains. Genomic data can provide important insights into the dissemination of bacterial subclades of global concern. The widespread AMR genotypes reported here highlight the need for improved diagnostics and updated treatment guidelines for bacteraemia in Ugandan cancer patients.
Neil, M.; Evans, B. A.
Show abstract
Multilocus sequence typing (MLST) remains the predominant method for typing bacterial strains. A common method for investigating particularly successful epidemic lineages within a species is to cluster isolates with similar MLST profiles into clonal complexes (CCs). Some CCs, such as international clones (ICs) in A. baumannii, are identified with specific sequence types (STs) and are of particular importance to human health. Although theoretically simple, there is a lack of convenient, user-friendly tools perform this analysis. Here we present PhyloMLST, a tool to cluster bacterial isolates into CCs and map them to ICs using the output from existing MLST tools and user-provided STs. As there is potential for aberrant IC assignment arising from excessively large CCs constructed with spurious links, PhyloMLST provides additional functionality to correct IC assignment with a user-provided phylogenetic tree. Although designed with A. baumannii in mind, PhyloMLST can be applied to any bacteria where construction and investigation of CCs based on MLST is performed. Impact statementMany bacterial pathogens are characterised by successful epidemic lineages that are responsible for a substantial number of infections, may be more virulent, and may carry an abundance of antimicrobial resistance genes. These epidemic lineages are comprised of a number of multilocus sequence typing (MLST) sequence types (STs), clustered into clonal complexes (CCs). To date, identifying which STs belong to which epidemic lineage has been challenging, with no simple analytical tools available. Here, we present PhyloMLST - a phylogenetically-aware method for assigning STs to epidemic lineages. The customisable nature of the tool will enable researchers to straightforwardly characterise any population of bacteria that they are working on using MLST data and user-defined definitions of epidemic lineages. Data summaryThe PhyloMLST source code and example data shown here is available at https://github.com/Mattn286/PhyloMLST.
Kenyon, J. J.
Show abstract
Acinetobacter baumannii is one of the most critical bacterial pathogens requiring novel approaches for infection control and treatment to curb the spread of highly resistant isolates. Epidemiological surveillance, and the design and application of many non-antibiotic interventions, require rapid and accurate prediction of diverse capsular polysaccharide (CPS) types from whole genome sequence data. The internationally adopted CPS typing system relies on the in-silico detection of CPS biosynthesis genes both in and outside the K locus (KL), utilising a reference sequence database that is compatible with the bioinformatics tool, Kaptive. In this study, 168 novel loci were added to the database following an extensive survey of publicly available A. baumannii genomes and non-redundant sequence entries, bringing the total number of reference sequences to 409 KL and 10 extra-locus genes. All novel loci conformed to the characteristic K locus configuration described previously, with conserved core genes for CPS biosynthesis flanking a central and highly variable region that includes type-specific structural genes. Across the 409 KL, there were 1000 protein clusters. Annotations for 309 novel clusters were curated and validated using a variety of sequence-and structure-based approaches. Most proteins (n=781) were encoded by genes found in [≤]4 KL, consistent with extensive structural diversity of CPS types in A. baumannii as reported previously. However, gene distribution analysis predicted shared structural features, with several KL predicting the same CPS type or K unit. Validation of the updated database against 46,185 publicly available genomes revealed that 32 K loci were present in 87.1% of sequenced isolates, with an overrepresentation of genomes with KL2, KL3, KL18, KL9 or KL22. IMPACT STATEMENTCapsular polysaccharide (CPS) typing is embedded in in-silico genomic approaches utilised for characterisation and epidemiological surveillance of Acinetobacter baumannii isolates. As new information regarding CPS diversity and structures becomes available, continual updates to the integrated typing database are essential to ensure accuracy and endurance as a valuable scientific resource. The rapid expansion of genome data released into public databases, particularly from previously underrepresented geographies and isolation sources in recent years, provided an opportunity to capture further diversity in CPS encoding regions to enhance the utility of the typing database. A total of 168 novel K loci were identified and added to the database, and an analysis of gene distribution across all 409 loci predicted shared CPS features. This work is expected to enhance KL typing and prediction of CPS structural features for studies involving diagnostics, surveillance and therapeutics applications. DATA AVAILABILITYAll genome sequences used in this study are publicly available with details and accession numbers listed in Supplementary Table 1. The updated database with 409 KL and 10 extra-locus reference sequences is freely available at https://github.com/johannajkenyon/Abaumannii_surface_polysaccharide_loci.
Lee, T. S. E.; Nguyen, L.; Forde, B. M.; Maidment, T.; Ye, S.; Henderson, A.; Playford, E. G.; Runnegar, N.; Henderson, B.; Watson, C.; Lindsay, M.; Bursle, E.; Douglas, J.; Hume, J.; Paterson, D. L.; Kidd, T.; Graves, B.; Hume, A.; Hall, M. B.; Schembri, M. A.; Beatson, S. A.; Harris, P. N. A.; Roberts, L. W.
Show abstract
OXA-48-like carbapenemases have been historically rare, however steady increases both locally and globally have warranted further investigation into their spread. Here we present the largest genomic analysis of blaOXA-181-producing bacteria in Australia to date, focusing on a single jurisdiction over seven years (2017 -- 2024). The initial investigation was prompted by an outbreak in 2017, where enhanced genomic surveillance in a single hospital identified 85 outbreak isolates related to an imported Escherichia coli ST38, carrying blaOXA-181 on an IncX3/colKP3 plasmid (previously reported as pOXA181). After four months of intensive infection control, the initial outbreak strain was eliminated. To confirm the outbreak plasmid was also contained, we collected all blaOXA-181-positive isolates from the same jurisdiction over subsequent years and sequenced with both Illumina and Oxford Nanopore Technologies to investigate clonal and mobile genetic element mediated spread. While continued surveillance post-2017 did not identify the same E. coli strain following the outbreak, pOXA181 plasmids were identified in >70% of surveillance isolates, with minimal genetic changes, which initially suggested local plasmid-mediated spread. Additional comparison to a global collection of pOXA181 plasmids found that epidemiologically unrelated pOXA181 plasmids were near identical, with no rearrangements and low, or no, single nucleotide polymorphisms. This suggests the mutation rate of pOXA-181 is incompatible with recent genomic transmission inference. This study highlights the current genomic epidemiology and drivers of blaOXA-181 and further demonstrates the necessity for detailed understanding of plasmid evolutionary rates to inform genomic surveillance.
Kassabian, L.; Al Khoury, C.; F Araj, G.; Tokajian, S.
Show abstract
Carbapenem-resistant Escherichia coli (CREc) recovered sequentially from one patient typically retain the same carbapenemase, with escalating resistance usually attributed to porin loss combined with pre-existing {beta}-lactamase expression. We used whole-genome sequencing to characterize a clonal pair of CREc isolates, CAEC145 and CAEC155, recovered 25 days apart from a hospitalized patient with sequential urinary and bloodstream infection. Both belonged to sequence type 361 (ST361), phylogroup A, serotype O-nontypeable:H30, and were separated by only 28 core-genome SNPs, confirming clonal relatedness. Despite this, the isolates differed sharply in carbapenemase content. CAEC145 carried blaOXA-1207, a recently described OXA-48-family variant, on a conjugative IncFII(pCoo)/ColKP3 plasmid, whereas CAEC155 lacked this determinant and instead harbored blaNDM-4 on a conserved IncX3 plasmid nearly identical to pJEG027, a member of a globally disseminated IncX3 lineage. This genotypic shift tracked a clear phenotypic transition. CAEC145 remained susceptible to imipenem and meropenem while resistant to ertapenem, whereas CAEC155 showed uniform high-level resistance to all three carbapenems and to ceftazidime-avibactam. Both isolates, however, remained susceptible to imipenem-relebactam, meropenem-vaborbactam, and cefiderocol. Comparative genomics linked the blaOXA-1207 element to a {Delta}Tn6361 transposon structure also found in the original German isolates where blaOXA-1207 was first described, and a 137-genome core-genome phylogeny placed both isolates within a globally disseminated ST361 lineage carrying multiple carbapenemase classes. These findings document, to our knowledge, the first within-host succession from an OXA-48-like to an NDM-type carbapenemase during a single sequential E. coli infection, driven by plasmid-level displacement rather than in-place gene evolution, with implications for genomic surveillance and antibiotic selection.
Colombi, E.; Ghaly, T. M.; Samarakoon, N.; Rajabal, V.; Tetu, S. G.
Show abstract
Horizontal gene transfer mediated by mobile genetic elements (MGEs) is a major driver of bacterial evolution and ecological adaptation. In the plant-associated genus Xanthomonas, multiple MGEs have been implicated in virulence, host specialisation, and environmental persistence, yet MGE diversity and evolutionary dynamics across the genus remain poorly understood. Here, we performed a comparative analysis of conjugative and mobilisable plasmids, integrative and conjugative elements (ICEs), integrative and mobilisable elements (IMEs), and their cargo genes across 516 complete genomes of three major Xanthomonas species: X. campestris, X. cissicola, and X. oryzae. We identified pronounced interspecific differences, with X. cissicola and X. campestris harbouring large and diverse MGE repertoires, comprising 28.3% and 26.7% of their respective pangenomes, whereas X. oryzae contained far fewer MGEs, making up only 3.6% of the identified pangenome. These differences were associated with host defence systems, including CRISPR-Cas and restriction-modification systems, and with variation in CRISPR spacer diversity. IMEs were the most abundant MGEs across all species, encoding diverse defence systems and accessory genes. ICEs exhibited signatures of horizontal transfer within and between species, and across genera. Notably, nearly identical ICEs carrying heavy-metal resistance genes were identified in Xanthomonas and Pseudomonas aeruginosa, indicating recent transfer between genera. MGEs collectively carried genes involved in virulence, interbacterial interactions, defence against phages, and plant cell wall degradation, with several elements associated with specific pathovars. Together, our findings establish MGEs as key drivers of genome plasticity and adaptive evolution in Xanthomonas, shaped by a dynamic interplay with host defence systems.
Saniya, S.; Khan, A. A.
Show abstract
BackgroundProbiotic bacteria occupy the same gut niches as enteric pathogens, prompting concern that probiotic strains might carry or contribute mobile antibiotic-resistance genes (ARGs). Genome-based screening is routinely used to assess this risk, but many screens use chromosome-level assemblies that may omit plasmid-borne, high-mobility cargo. We quantified this effect and compared the mobile context resistomes of probiotic-associated and pathogen reference genomes. MethodsWe screened 50 bacterial reference genomes (25 probiotic-associated, 25 pathogen/comparator) using a reproducible workflow with the CARD nucleotide catalog, PlasmidFinder replicons, and ISfinder insertion sequences. Each ARG was assigned a fourtier in silico mobile-context risk category from plasmid co-localization and insertion-sequence (IS) flanking. The identical strain panel was screened in matched full-assembly and chromosome-only modes. Acquired calls were curated against intrinsic/efflux/biocide determinants and cross-checked with AMRFinderPlus, ResFinder, and targeted BLAST, with MOB-suite as a plasmid/mobility overlay. ResultsFull-assembly screening detected 373 ARG loci versus 338 in chromosome-only mode on the same strains, increasing High-risk calls from 4 to 15 and recovering 32 plasmid replicons (chromosome-only: 0). All 15 High-risk mobile-context loci occurred in pathogen/comparator genomes and none in probiotic-associated genomes; no ARG was shared across groups at [≥]95% nucleotide identity (0/175 edges). Per-strain ARG burden was higher in pathogen genomes (mean 12.32 versus 0.52 loci; Mann-Whitney U = 606.5, P< 0.001). Most priority High-risk loci were corroborated by one or more external tools, with discordant calls retained explicitly as flagged records. ConclusionsChromosome-only screening materially undercounts mobile ARG cargo. In this reference-genome panel, high-risk mobile-context loci were concentrated in pathogen/comparator genomes -- an in silico reference-genome-level safety signal rather than evidence about commercial products or genetic transfer. Data SummaryNo new sequencing data were generated; all genomes are publicly available reference assemblies from NCBI RefSeq. O_LIGenome accessions for all 50 genomes are in Supplementary Table S1. C_LIO_LISource code (screening, validation, and figure scripts, including make_figures.py) is available at https://github.com/abdullahak07/1dprob. C_LIO_LIProcessed result tables are summarised in Supplementary Tables S2-S10. C_LIO_LIExternal validation outputs (AMRFinderPlus, ResFinder, BLAST, MOB-suite) are provided as supplementary data. C_LIO_LITool/database versions and detection thresholds are in Supplementary Table S9. The authors confirm that all supporting data, code, and protocols are available within the article, the cited repositories, or the supplementary material. C_LI O_TBL View this table: org.highwire.dtl.DTLVardef@b0f2beorg.highwire.dtl.DTLVardef@110c676org.highwire.dtl.DTLVardef@5578ceorg.highwire.dtl.DTLVardef@16e448forg.highwire.dtl.DTLVardef@573000_HPS_FORMAT_FIGEXP M_TBL O_FLOATNOSupplementary Table S9:C_FLOATNO O_TABLECAPTIONtool/database versions and detection thresholds used in the SCALE50 analysis. Database snapshots and external validation outputs should be archived with the supplementary data to ensure reproducibility. C_TABLECAPTION C_TBL Impact StatementGenome-based screening is widely used to judge whether a bacterial strain carries transferable antibiotic-resistance genes, including in the safety assessment of probiotic-associated species. We show that the assembly level chosen for such screening materially affects the result: on an identical panel of 50 reference genomes, chromosome-only analysis recovered fewer than a third of the high-risk, mobile-context resistance loci detected when plasmid replicons were included. Applying full-assembly screening with transparent, multi-tool external validation, high-risk mobile-context resistance genes were concentrated in pathogen/comparator genomes and absent from the probiotic-associated genomes in this panel, with no cross-group sharing at high identity. These observations argue for full-assembly inputs and explicit mobile-context interpretation in genome-based resistance screening and provide a cautious, reference-genome-level safety signal; they are not claims about commercial products and do not demonstrate genetic transfer.
Kulkarni, S. M.; Jacob, J. J.; Rajendra, S.; S, P.; T, M. P.; Velmurugan, A.; Nelson, R.; Neeravi, A.; Balaji, L.; Gunasekaran, K.; Manesh, A.; Rajni, E.; Walia, K.; Veeraraghavan, B.
Show abstract
Carbapenem-resistant Klebsiella pneumoniae (CRKp) is a critical global healthcare threat driven by high-risk multidrug-resistant (MDR) clones that acquire hypervirulence genes. Although resistance-virulence co-occurrence is extensively documented, the plasmid-level mechanisms facilitating this convergence remain unclear. In this study, we utilized hybrid short- and long-read whole-genome sequencing of 376 clinical CRKp strains to define the evolutionary trajectories and structural plasmid dynamics of three predominant high-risk clones: ST147 (n=157), ST231 (n=108), and ST2096 (n=111). Carbapenemase genes were present in 90% of isolates, predominantly blaOXA-48-like and blaNDM-5 co-harbored with blaCTX-M-15. Virulence profiling indicated high aerobactin (iuc) prevalence (62.7%), while salmochelin and colibactin were undetected. Hypermucoviscosity occurred infrequently (6.6%) and was independent of rmpA/rmpA2, confirming a clear genotype-phenotype discordance. Comparative plasmid mapping revealed three distinct, lineage-specific plasmid configurations underlying this intermediate convergent pathotype: ST147 exhibited dynamic, mosaic hybrid IncFIB-IncHI1B plasmids; ST2096 showed structurally stabilized hybrids; and ST231 retained virulence and resistance determinants on separate, segregated plasmids. These findings show that convergence is regulated by multiple, clone-specific evolutionary routes rather than a single path, highlighting the critical need for more in-depth genomic surveillance capable of identifying convergent plasmids along with high-risk lineages
Hasugian, I. A.; Alifiyah, N. I.
Show abstract
Background/aimAntimicrobial resistance in methicillin-resistant Staphylococcus aureus (MRSA) requires precision non-antibiotic therapeutics, yet phage lytic efficacy is poorly predicted by phenotypic assays, as shown by paradoxical biofilm responses. This study characterized the genomic architecture of lytic S. aureus bacteriophages, focusing on the conservation of the lysis module and the variability of host-recognition modules, to provide a rational basis for phage candidate selection. Materials and methodsTwenty-two complete S. aureus phage genomes were retrieved from NCBI GenBank. Genomic features were extracted with custom Biopython scripts. Lysis (endolysin, holin) and host-recognition (tail fiber/receptor-binding protein) modules were annotated and validated by InterPro domain analysis, with disrupted endolysins resolved by tBLASTn. Phylogeny was reconstructed from large terminase subunit (TerL) sequences using maximum likelihood. ResultsGenome size spanned three classes, from 17.5 to 148.6 kb. The LysK-type endolysin (CHAP-Amidase-SH3b) was highly conserved, whereas tail fiber/RBP genes were detected in only 14 of 22 phages. Domain analysis reclassified two proteins annotated as endolysins as virion-associated peptidoglycan hydrolases, and identified two independent mechanisms--HNH endonuclease insertion and intron splitting--that interrupt lysis-module genes and confound automated annotation. Maximum likelihood analysis recovered a strongly supported, highly conserved core clade with EW and SA13 as divergent lineages. ConclusionLysis modules are conserved whereas host-recognition modules are variable, indicating that host recognition rather than the lytic enzyme is the principal determinant of host range and the more rational target for phage selection and engineering.
Ussery, D.; Bukharid, M. Z.; Majumder, R.; Borin, V. A.; Alisoltani, A.
Show abstract
Prokaryotic taxonomy now relies on both marker-gene and genome-wide sequence comparisons, but these methods differ in taxonomic range, scalability, and sensitivity to genome quality. Here, we benchmarked four commonly used approaches, including 16S rRNA identity, FastANI, Mash distance, and FastAAI/Jaccard similarity across a dataset of 30,495 prokaryotic type-strain genomes. Type-strain genomes provide nomenclatural anchors for validly named species, making them a useful framework for evaluating how sequence-based methods correspond to current taxonomic assignments. We evaluated method behavior across taxonomic ranks from species to domain and separated initial method failures from threshold-based failures. When clean full-length 16S rRNA sequences were available, same-species comparisons passed the empirical threshold in >97% of cases. However, a usable full-length 16S rRNA sequence was unavailable for 4,551 of the 16,402 same-species comparisons (28%), limiting marker-gene-based analysis. In addition, 16S rRNA identity ranges overlapped across higher taxonomic ranks, limiting the use of universal rank-specific cutoffs. FastANI provided strong species-level resolution, with same-species comparisons passing the empirical threshold in approximately 88% of cases but was less informative at deeper ranks. Mash enabled rapid genome-scale screening, although its distance values require careful interpretation beyond close relatives. FastAAI provided a genome-wide amino-acid signal, with approximately 92% of same-species comparisons passing the empirical threshold and was especially useful for comparisons beyond the species boundary. Overall, no single method performed optimally across all taxonomic levels. These results support a rank-aware benchmarking framework in which 16S rRNA, FastANI, Mash, and FastAAI are interpreted as complementary tools, with attention to genome quality, missing data, and method-specific failure modes.
Ounissi, N. E.; Gomri, M. A.; El Hadef El Okki, M.
Show abstract
The identification of novel probiotic candidates with potential health-promoting properties remains a major challenge in food biotechnology and increasingly relies on in silico screening of genomic information. However, probiogenomic markers are heterogeneous, and safety, survival-colonisation, and functional-benefit traits do not contribute equally to probiotic potential. This study developed the Structured Probiotic Potential Index (SPPI), a fuzzy multicriteria system for genome-based probiotic candidate prioritisation. A hierarchical evaluation structure was established from probiogenomic evidence and organised into three main pillars and fourteen subcriteria. Expert judgements were collected using the Analytic Hierarchy Process, followed by consistency-based curation and weight aggregation. A curated dataset of 48 complete bacterial genomes, distributed into probiotic, potentially probiotic, neutral, and pathogenic groups, was taxonomically validated and analysed using an automated probiogenomic screening pipeline. Genome-wide screening generated 3,218 binary genomic features, from which curated probiogenomic markers were mapped to the scoring hierarchy. The resulting index was formulated as a normalised expert-weighted equation integrating safety, survival-colonisation, and functional-benefit components. SPPI prioritised genomes according to weighted probiogenomic profiles and separated pathogenic genomes from favourable probiotic and potentially probiotic profiles within the analysed dataset. Downstream Fuzzy Comprehensive Evaluation transformed the continuous score into probiotic/potentially probiotic, neutral, and pathogenic classes. Under internal leave-one-out evaluation, all 48 genomes were assigned to their expected reference classes, while confidence analysis distinguished high-confidence from borderline assignments. Post-classification comparison with ProbML showed concordant behaviour for most genomes and discordant predictions for selected cases. These results support SPPI as a transparent genome-based decision-support system for early probiotic candidate prioritisation before experimental validation.
Elena, A. X.; Batantou Mabandza, D.; Kluemper, U.; Breurec, S.; Dagot, C.; Berendonk, T. U.
Show abstract
The global dissemination of antimicrobial resistance is increasingly driven by bacterial clones combining antimicrobial resistance with enhanced virulence and environmental adaptability. Escherichia coli sequence type 131 (ST131) has historically been regarded as a major disseminator of the extended-spectrum {beta}-lactamase (ESBL) blaCTX-M-15. However, the emergence of E. coli ST1193 carrying blaCTX-M-15 may represent an ongoing shift in the epidemiology of this resistance determinant. Here, we investigated the prevalence, genomic characteristics, virulence and antimicrobial resistance potential of ST1193 in comparison with ST131. A total of 1,136 E. coli isolates were recovered from touristic and non-touristic environments, hospital-associated samples, and aircraft toilets in Guadeloupe. Isolates were whole-genome sequenced and analysed for antimicrobial resistance and virulence determinants. Additionally, publicly available genomic data comprising 1,215 blaCTX-M-15-positive ST131 and ST1193 isolates were analysed to assess temporal and geographical trends. ST1193 was significantly associated with aircraft-associated samples and exhibited a higher antimicrobial resistance gene burden than ST131, while maintaining a comparable virulence factor content. Analysis of publicly available genomes revealed similar temporal emergence patterns for blaCTX-M-15-positive ST1193 and ST131, with ST1193 showing a more recent distribution and a higher number of deposited isolates in recent years, consistent with a potential ongoing clonal replacement. Comparative genomic analysis identified numerous virulence and adaptation-associated genes shared between both sequence types, while ST1193 additionally carried distinct determinants, including components of the transmissible locus of stress tolerance. Furthermore, quinolone resistance-associated mutations were strongly linked to blaCTX-M-15 carriage, particularly among ST1193 isolates. Together, these findings identify E. coli ST1193 as an emerging high-risk clone with substantial potential for blaCTX-M-15 dissemination. Its association with aircraft-associated samples further highlights the potential role of air travel in long-distance transmission and underscores the need to reconsider current surveillance strategies focused predominantly on ST131.
Morin, K.; Hetrick, E.; Shrivastava, A.; Tran, E.; Cartee, J. C.; Hebrank, K.; Gernert, K.; Schmerer, M.; Joseph, S. J.
Show abstract
Antimicrobial-resistant Neisseria gonorrhoeae (Ng) poses a growing global public health threat. Existing tools available for Ng genome analysis carry notable limitations, including incomplete resistance marker coverage, absence of species identification, and lack of phylogenetic capability. We developed neiss-flow, a highly parallelized Nextflow pipeline for Ng genome analysis that integrates five subworkflows: read preprocessing, species identification, de novo assembly, antimicrobial resistance (AMR) profiling, and recombination-aware phylogenetic analysis with outbreak detection. neissflow performs extensive quality control on reads, assemblies, variant calls, and phylogenetic results. Validation was performed on two datasets: a mixed-species dataset (n=158; 105 Ng and 53 non-gonococcal species) for sensitivity/specificity assessment, and a reproducibility dataset (n=283 replicate sequences from 17 reference strains) for consistency and phylogenetic validation. neissflow achieved 100% sensitivity and specificity for Ng species identification compared with MALDI-TOF and PubMLST methods. All nine AMR and typing analytes demonstrated [≥]98.1% concordance with PubMLST genotype calls. Genotype-phenotype validation confirmed perfect concordance for key resistance determinants including gyrA mutations with ciprofloxacin resistance, 23S rRNA mutations with high-level azithromycin resistance, and tetM plasmid gene with high-level tetracycline resistance. Reproducibility analysis demonstrated 99.97% concordance across 3,093 analyte calls. Phylogenetic validation demonstrated 100% accuracy for both strain-level and intra-MLST clustering. neissflow is a robust, accessible, and standardized pipeline, positioning it as a valuable tool for public health laboratories engaged in Ng AMR monitoring and outbreak investigations. ImportanceNeisseria gonorrhoeae (Ng) is the second most common reported bacterial sexually transmitted infection and has developed resistance to all clinically relevant antibiotics. Surveillance of Ng resistance informs clinical recommendations for treatment of gonococcal infections. Whole genome sequencing (WGS) offers powerful insights into resistance mechanisms and transmission dynamics. However, many public health laboratories lack the resources needed to analyze these data effectively. We developed neissflow as an end-to-end, accessible Ng WGS analysis pipeline. neissflow demonstrated exceptional accuracy and reproducibility across diverse reference datasets. neissflow enables broader adoption of whole-genome sequencing-based surveillance and supports timely public health responses to emerging antimicrobial resistance in Ng by lowering technical barriers.
Farid, A. C.; Haldeman, S.; Otto, C.; DMello, A.; Tettelin, H.; Ratner, A. J.
Show abstract
Based on recent epidemiologic studies, Streptococcus agalactiae (Group B Streptococcus; GBS) sequence type (ST) 1010 is an emerging lineage now identified in multiple countries. We report the phylogenetic and genomic characteristics of a set of 55 GBS sequence type (ST) 1010 strains, as well as two newly described single-locus variants of ST1010. A core genome phylogeny suggests that ST1010 is closely related to both ST452 and the hypervirulent clonal complex (CC) 17 GBS lineage. Notably, we demonstrate that genes encoding two virulence determinants previously described as specific to CC17 GBS, the HvgA adhesin and the serine-rich repeat protein Srr2, are both present in ST1010 genomes. Srr2 is shared with members of ST452. High-level gentamicin resistance (HLGR) encoded on an IS256 mobile element, previously described in a small number of ST1010 isolates, is present in a distinct ST1010 subclade encompassing the majority of ST1010 isolates. The relationship between ST452 (serotype IV), ST1010 (serotype IV), and ST17 (serotype III) strains suggests that ST17 may have arisen from a serotype IV ancestor and later acquired the type III capsule locus. Taken together, these findings clarify the phylogenetic position of ST1010 and suggest sequential acquisition of virulence determinants and HLGR prior to its international emergence. IMPACT STATEMENTST1010 GBS has emerged internationally, with colonizing and invasive isolates described in the United States, Dominican Republic, Netherlands, and Italy. Using a core genome phylogeny and targeted detection of genomic regions, we demonstrate that ST1010 shares specific virulence determinants with the CC17 hypervirulent GBS lineage and that HLGR is confined to a specific numerically dominant subclade of ST1010. Our work spotlights the importance of future epidemiologic and genomic surveillance of ST1010 and related lineages. DATA SUMMARYPublicly available genomic data were used from three previously published studies (Laycock KM et al., McGee L et al., Khan UB et al.), as well as a set of newly sequenced GBS genomes from clinical strains originating in New York City (NYC). The corresponding accession numbers and detailed information for all strains are provided in the Table.
Allam, C.; Charmat, Y.; Agsous, S.; Awad, Z.; Fouchet, T.; Goncalves, L.; Ben Salem, N.; Poignon, C.; Mougari, F.; Veziris, N.; Cambau, E.
Show abstract
Macrolides are key agents for treating infections caused by non-tuberculous mycobacteria (NTM). Nevertheless, chromosomal erm genes conferring inducible macrolide resistance are described in some NTM species, such as Mycobacterium abscessus and M. fortuitum, whereas M. chelonae had long been considered as lacking a functional erm. Recent descriptions from the USA and Japan of a new plasmid-borne erm(55) (erm(55)P) in M. chelonae and other rapidly growing mycobacteria (RGM) have challenged this assumption. We investigated erm(55)P occurrence in clinical RGM referred to the French National Reference Centre for Mycobacteria between 2012 and 2026 by genome screening and erm(55)P specific real-time PCR. Positive isolates underwent long-read whole genome sequencing (GridIon, Oxford Nanopore Technologies). Clarithromycin (CLR) minimum inhibitory concentration (MIC) was determined by broth microdilution (RAPMYCO and FRATMYC, Thermo Fisher) and read up to 14 days. Five clinical isolates showing inducible CLR resistance (MIC range <0.25-64 mg/L on day 3-4 and 128 - >128 mg/L on day 14) were positive for erm(55)P: one M. chelonae, three M. neoaurum, and one M. parafortuitum. erm(55)P-positive M. chelonae genomes from this and previous descriptions did not cluster together in the phylogenetic analysis of 263 genomes. The assembled plasmids showed high similarity to previously reported erm(55)-carrying plasmids, especially within the erm(55)P region. The upstream sequence of erm(55)P showed a secondary structure compatible with a possible translation attenuation mechanism. These findings document the first report of a plasmid-borne erm(55) in Europe in M. chelonae and other RGM and raise concern about the emergence of plasmid macrolide resistance in NTM.